Papers with uniform sparsity rates
Weight-Aware Activation Sparsity with Constrained Bayesian Optimization Scheduling for Large Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing activation sparsification methods rely on activation magnitude and weights for sparsity . authors propose a weight-aware activation-a-ware framework for large language models . |
| Approach: | They propose a weight-aware activation sparsity framework that uses weight-based scoring to measure activation importance in sparsification and a custom GPU sparse kernel to support it. |
| Outcome: | The proposed framework outperforms existing methods at 60% model-level sparsity and significantly outperfies them at higher sparsities. |